fix(windows/capture): production hardening from a real 3-project deployment (closes #1, closes #2) + opt-in --global single-daemon mode - #3
Conversation
…ngle-daemon mode - zcode detectCompaction: reuse cached getConn (was new Database per session per cycle, never closed - handle leak + memory sawtooth root cause, upstream EwanJasper#2); healthCheck closes probe conn - cli/service windowsRunScript: quote nodeAbs (upstream EwanJasper#1) + chcp 65001 for CJK paths; vbs written UTF-16LE BOM (wscript misreads UTF-8 CJK paths as GBK) - capture/sync: per-session content-signature watermark (adapter.changeProbe: zcode MAX(rowid) aggregate, 2 SQL/cycle; file sources mtime+size) + known-in-db guard (rebuild/forget safe) + 10-min full sweep backstop -> idle daemon CPU 7.5% -> 0.08% - capture/watch: per-project named-pipe single-instance lock (upstream EwanJasper#2 instance accumulation); worker extraction; --global mode: one process adopts all registry projects dynamically
…d --global process was adopting 0 projects and idling as zombie
… projects were unreleasable (sqlite/cmd held open by this process), now auto-pruned within one sweep
…obal lock released on exit); fix ensureRegistered double-call that swallowed custom-adapter load errors (upstream wart)
…决策'); upstream wart exposed by brief quality audit
…markdown fragments / embedded JSON residue) + unresolved 14-day fade in CLI
… opts.json - upstream wart the brief's text-fallback was papering over)
- mcp adopt(): close previous project's db on switch (handle leak + EBUSY vs forget/rebuild) - archive: days/sizeMb 0 no longer treated as unset (0-days archive-everything footgun) - loadConfig: corrupt config.json -> readable error instead of raw stack - forget: corrupt preview snapshot now REJECTS --yes (fail-safe optimistic lock) - annotate_session/save --tag: re-include decision texts in meta_text (viaMeta search surface)
…assertion broke past the 30-day boundary on Sep 19) - suite is fully green (222) for the first time
…h.log - per-project logs froze forever under the global daemon, making 'watch --status' print zombie tails in non-cwd projects
- watcher: runtime 'error' events (dir deleted / overflow / AV) had no listener -> EventEmitter throw kills the whole global daemon; now degrades to the existing 1s polling fallback (self-heal) - insertMessage: index text capped at 100KB (jieba sync tokenize of huge tool-dump messages stalled the event loop seconds); full content still stored, only FTS coverage capped - zcode discover: cache by source-db mtime (idle cycles = 0 SQL instead of unindexed lower() table scan per worker per event) - watch: wire the local anonymous stats counter into the daemon cycle (status panel's 'ignored 0' lied vs 34 real blocks)
… with N workers sharing one source db, projects B/C received project A's session list and ingested it as their own (cross-project contamination, live: music 42->559, novel 15->495). Key now includes project root.
…n (4th-project field report) From a live field report (E:/mc整合包, 4th deployment, direct SQLite forensics): - P0 decisions froze for 20+h on long-lived sessions: extraction only ran at confirm (idle 10min -> pending -> 6h cooldown -> confirm); continuously-worked sessions never completed the loop. Fix: refreshExtraction() now runs at every pending transition (10-min idle), decisions lag <= 10min. confirmSession reuses it. - P1 88% of extracted decisions were workflow-subagent fragments: zcode adapter now tags title-matched subagent sessions with originHint='workflow' -> origin column -> listDecisions excludes them by default (data preserved, queries clean). Live: 36 sessions auto-tagged, visible decisions 82 -> 20, residue 0. - P2 decision text hard-cut at 90 chars produced mid-sentence/trailing-bracket fragments: cut at nearest pause char + trim trailing closers + reject leading-closer fragments. Tests: 222 -> 221 pass + 1 skipped, 0 failures.
…text (same family as annotate/save fix - confirm/refresh path was the missed third site)
…n captured twice displayed twice in briefs)
Follow-up: 3 more commits since opening (live field report from a 4th deployment)A 4th project ( P0 — decisions froze for 20+ hours on long-lived sessions. Extraction only ran at the confirm transition (idle 10min → pending → 6h cooldown → confirm); a continuously-worked session never completes that loop, so P1 — 88% of expanded decisions were workflow-subagent fragments. The zcode adapter now tags title-matched subagent sessions with P2 — decision text hard-cut at 90 chars produced mid-sentence / trailing-bracket fragments: cut at the nearest pause char + trim trailing closers + reject leading-closer fragments. Plus: Commits: 3e7679e (P0-P2), f17f52d (user_tags), 3b69246 (dedupe). Suite still 222 → 221 pass + 1 skipped, 0 failures. All changes verified on the live 4-project deployment (heartbeats, capture latency, brief contents). |
…ntence start; 最终用/最终选 stay strong (they are genuine final-choice verbs - caught by the suite)
…rmark from decisions column, O(delta) instead of O(full session)); topics/summary full rebuild stays at confirm - kills the last recurring CPU/RAM burst generator on long-lived sessions
…le/bold residue rejected); briefs read unresolved entries via q field (title-only display was a silent regression from the --json fix)
…de confirm counts were invisible in the status panel)
… and incremental decision refresh (seq-watermark idempotency); guards the mc-report root-cures against silent breakage
Context
This branch came from running SessionRelay 0.5.1 as the memory layer for three real projects on one heavy Windows machine (i7-14700HX, 3 project roots sharing one ZCode source db, ~25k captured messages). Two red-team rounds + a full source audit surfaced the root causes behind #1 and #2. Everything below is verified against your own test suite: 222 tests → 221 pass + 1 skipped, 0 failures (also defused a time-bomb in
test/forget/functional.spec.tswhose absolute fixture date broke its own assertion past a 30-day boundary — the suite is fully green for the first time).Fixes
#1 — Windows boot chain (
src/cli/service.ts)windowsRunScript: quotenodeAbs(C:\Program Files\nodejscontains spaces →'C:\Program' is not recognized), and emitchcp 65001 >nulfirst — cmd.exe parses the file in the OEM codepage (GBK on Chinese systems) while we write UTF-8, so CJK project paths corruptedcd /dand the log redirect.#2 — daemon resource/lifecycle (
src/capture/watch.ts,src/cli/watch.ts,src/adapters/zcode/index.ts)detectCompactionopened a new better-sqlite3 connection per session per cycle and never closed it (thefinallywas empty — a stale "cache connection" comment copied fromgetConn). Now reusesgetConn.healthCheckcloses its probe too. Production evidence: handle slope ~378/h → flat.\\.\pipe\srelay-watch-<sha1(root)[:16]>) — kernel-owned, dies with the process, EADDRINUSE = already running. win32-only; other platforms unchanged.Capture correctness / quality
extract.ts: decision extraction matched trigger words anywhere in a sentence and captured kanban table rows, markdown bold fragments and embedded-JSON residue. Added a plausibility guard (min length, trailing colon/pipe/bold rejection, table-row rejection, JSON residue rejection).bin/srelay.tsunresolved:--jsonexisted but was dropped in the action — always emitted human text. Now passed through.cli/meta.tsdecisions --limit: took the oldest N (ascending sort +slice(0, n)) while labelled "recent" — now takes the newest N;unresolvedgets a 14-day fade.capture/sync.tsrunSync: added a per-session content-signature watermark (changeProbe: zcode = two aggregate MAX(rowid) queries per cycle; file sources = mtime+size) plus a "session still in db" guard (so rebuild/forget flows still re-ingest) and a 10-minute full-sweep backstop. Idle daemon CPU went from ~7.5% of a core to ~0.1%.store/db.tsinsertMessage: FTS index text capped at 100KB (jieba sync tokenize of multi-hundred-KB tool dumps stalled the event loop; full content is still stored).Other warts found in the same audit
capture/sync.ts:ensureRegisteredcalled twice (second call returns empty — custom adapter errors were never reported).capture/archive.ts:--days 0/--size 0mbtreated as unset (falsy) → cutoff degraded to 9999-12-31, "archive nothing" archived everything (with--hard).shared/config.ts: corruptconfig.jsonproduced a raw stack in every CLI command → readable error with repair hint.cli/forget.ts: corrupt preview snapshot silently skipped the optimistic-lock diff before--yes→ corrupt snapshot now refuses execution (fail-safe).annotate_session/save --tag: rewritingmeta_textdropped decision texts → viaMeta search surface silently narrowed; now re-included.adapters/claude-code/watcher.ts: runtimeerrorevents fromfs.watchhad no listener — a deleted/overflowed watched dir threw and killed the whole daemon; now degrades to the existing 1s polling fallback.Opt-in feature:
srelay watch --globalOne daemon process adopts all projects from the existing registry (
registryCandidates()), spawns per-project workers (each still holds its own per-project pipe/file locks), re-scans every 10 min to adopt new / prune removed projects, and dual-writes worker logs to each project's ownwatch.log(per-project--statustails stay fresh). Guards against a second global instance via a fixed-name pipe. Deployed on the 3-project machine: 3 daemons → 1, idle CPU ≈ 0, cross-project isolation verified.Notes
better-sqlite3connections all carrybusy_timeoutalready, so daemon+CLI+MCP concurrent writes are safe.--globalpart is independent of the fixes.Closes #1, closes #2.